Skip to content

[Klaud Cold] Update dsr1-fp8-h100-dynamo-sglang SGLang image to v0.5.19-cu130 / 将 dsr1-fp8-h100-dynamo-sglang 的 SGLang 镜像升级至 v0.5.19-cu130 - #2967

Closed
Klaud-Cold wants to merge 1 commit into
mainfrom
klaud/auto-e4b2c5440222f47a-30f4b66275d7e326
Closed

[Klaud Cold] Update dsr1-fp8-h100-dynamo-sglang SGLang image to v0.5.19-cu130 / 将 dsr1-fp8-h100-dynamo-sglang 的 SGLang 镜像升级至 v0.5.19-cu130#2967
Klaud-Cold wants to merge 1 commit into
mainfrom
klaud/auto-e4b2c5440222f47a-30f4b66275d7e326

Conversation

@Klaud-Cold

Copy link
Copy Markdown
Collaborator

Update the dsr1-fp8-h100-dynamo-sglang master image from lmsysorg/sglang:v0.5.8-cu130 to the current SGLang release lmsysorg/sglang:v0.5.19-cu130 (Docker Hub digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, tag commit sgl-project/sglang@0bcd822). Model, disaggregated 1P/1D TP16 topology, EAGLE/MTP settings, the two upstream srt-slurm recipe references and the 8k1k concurrency lists are unchanged.

Baseline

  • Published date: 2026-02-13 (benchmarks?model=DeepSeek-R1-0528&date=2026-02-13&exact=true, workflow-info?date=2026-02-13, evaluations?model=DeepSeek-R1-0528&date=2026-02-13&exact=true)
  • Old image: lmsysorg/sglang:v0.5.8-cu130 (Docker Hub digest sha256:ef0d14df76c2c90ce651c274bc607600d09426128231722a96d52bb5472e1ebf, tag commit sgl-project/sglang@0189f41). Baseline mismatch: the published rows carry this label, but runners/launch_h100-dgxc-slurm.sh maps every H100 dynamo-sglang job to the staged squash file lmsysorg_sglang_v0.5.8.post1-cu130.sqsh, so the producing container was lmsysorg/sglang:v0.5.8.post1-cu130 (digest sha256:b6f9f50829ec45428db4451616978c45eb3d8d91488a702e78bccd06154d6bef).
  • Workload / topology: H100 cluster:h100-dgxc, deepseek-ai/DeepSeek-R1-0528 FP8, Dynamo + SGLang disaggregated, NIXL KV transfer, EAGLE MTP (2 steps, top-k 1, 3 draft tokens), fixed-seq-len 8k1k (ISL 8192 / OSL 1024), two deployment shapes of 4 nodes each: 1 prefill worker TP16 EP1 plus 1 decode worker TP16 EP1 (upstream recipe recipes/h100/8k1k/mtp/h100-fp8-1p1d-max-tp-mtp.yaml, concurrency 1-128) and 1 prefill worker TP16 EP1 plus 1 decode worker TP16 EP16 with DP attention (recipes/h100/8k1k/mtp/h100-fp8-1p1d-max-dep-mtp.yaml, concurrency 1-64), both from NVIDIA/srt-slurm sa-submission-q2-2026
  • Producer: run 21973157671 (head 78d48618f771603c6a06e61b0083362814d30919, changelog PRs #643 and #644); 15 benchmark points, result IDs below

1P TP16 EP1 / 1D TP16 EP1 (TEP)

Conc Result ID Total tok/s/GPU Output tok/s/GPU Median TTFT (s) Median TPOT (ms) Median E2E (s)
1 94693 29.0 6.5 0.560 9.48 9.24
2 94698 51.1 11.5 0.600 9.99 9.68
4 94700 87.5 19.5 0.611 12.24 11.63
8 94702 127.2 28.6 0.625 16.87 16.04
16 94694 178.6 39.6 0.853 24.51 22.77
32 94699 230.3 51.5 0.987 37.62 35.71
64 94701 279.8 62.1 13.348 50.46 59.04
128 94692 279.7 62.0 72.637 48.76 116.48

1P TP16 EP1 / 1D TP16 EP16 DP-attention (DEP)

Conc Result ID Total tok/s/GPU Output tok/s/GPU Median TTFT (s) Median TPOT (ms) Median E2E (s)
1 94645 13.8 3.1 2.139 18.49 18.84
2 94646 27.1 6.1 0.744 19.47 19.30
4 94647 51.9 11.5 0.754 20.55 18.80
8 94649 92.9 20.9 0.875 22.63 21.37
16 94643 159.0 35.2 0.923 27.15 25.61
32 94644 243.2 54.4 1.022 34.71 32.89
64 94648 396.6 88.0 2.264 40.94 40.11
  • Published eval: N/A for 2026-02-13 (no evaluation rows ingested for this family on the baseline date). The nearest published gsm8k rows for the family are dated 2026-04-23 (run 24816108884, TEP concurrency 64 em_strict 0.9545 / em_flexible 0.9553, DEP concurrency 32 em_strict 0.9575 / em_flexible 0.9583) and are not part of the frozen baseline.

dsr1-fp8-h100-dynamo-sglang 的主配置镜像从 lmsysorg/sglang:v0.5.8-cu130 升级到当前 SGLang 发布版 lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,标签提交 sgl-project/sglang@0bcd822)。模型、分离式 1P/1D TP16 拓扑、EAGLE/MTP 设置、两个上游 srt-slurm 配方引用以及 8k1k 并发列表均保持不变。

基线

  • 发布日期: 2026-02-13(benchmarks?model=DeepSeek-R1-0528&date=2026-02-13&exact=trueworkflow-info?date=2026-02-13evaluations?model=DeepSeek-R1-0528&date=2026-02-13&exact=true
  • 旧镜像: lmsysorg/sglang:v0.5.8-cu130(Docker Hub 摘要 sha256:ef0d14df76c2c90ce651c274bc607600d09426128231722a96d52bb5472e1ebf,标签提交 sgl-project/sglang@0189f41)。基线不一致:已发布数据行标注的是该镜像,但 runners/launch_h100-dgxc-slurm.sh 将所有 H100 dynamo-sglang 任务映射到预置的 squash 文件 lmsysorg_sglang_v0.5.8.post1-cu130.sqsh,因此实际产出数据的容器是 lmsysorg/sglang:v0.5.8.post1-cu130(摘要 sha256:b6f9f50829ec45428db4451616978c45eb3d8d91488a702e78bccd06154d6bef)。
  • 工作负载 / 拓扑: H100 cluster:h100-dgxcdeepseek-ai/DeepSeek-R1-0528 FP8,Dynamo + SGLang 分离式部署,NIXL KV 传输,EAGLE MTP(2 步、top-k 1、3 个草稿 token),固定序列长度 8k1k(ISL 8192 / OSL 1024),两种各占 4 节点的部署形态:1 个 prefill worker TP16 EP1 加 1 个 decode worker TP16 EP1(上游配方 recipes/h100/8k1k/mtp/h100-fp8-1p1d-max-tp-mtp.yaml,并发 1-128),以及 1 个 prefill worker TP16 EP1 加 1 个 decode worker TP16 EP16 并启用 DP attention(recipes/h100/8k1k/mtp/h100-fp8-1p1d-max-dep-mtp.yaml,并发 1-64),均来自 NVIDIA/srt-slurm sa-submission-q2-2026
  • 数据来源: 运行 21973157671(head 78d48618f771603c6a06e61b0083362814d30919,changelog PR #643#644);共 15 个基准点,结果 ID 见下表

1P TP16 EP1 / 1D TP16 EP1(TEP)

并发 结果 ID 总吞吐 tok/s/GPU 输出吞吐 tok/s/GPU TTFT 中位数 (s) TPOT 中位数 (ms) 端到端中位数 (s)
1 94693 29.0 6.5 0.560 9.48 9.24
2 94698 51.1 11.5 0.600 9.99 9.68
4 94700 87.5 19.5 0.611 12.24 11.63
8 94702 127.2 28.6 0.625 16.87 16.04
16 94694 178.6 39.6 0.853 24.51 22.77
32 94699 230.3 51.5 0.987 37.62 35.71
64 94701 279.8 62.1 13.348 50.46 59.04
128 94692 279.7 62.0 72.637 48.76 116.48

1P TP16 EP1 / 1D TP16 EP16 DP-attention(DEP)

并发 结果 ID 总吞吐 tok/s/GPU 输出吞吐 tok/s/GPU TTFT 中位数 (s) TPOT 中位数 (ms) 端到端中位数 (s)
1 94645 13.8 3.1 2.139 18.49 18.84
2 94646 27.1 6.1 0.744 19.47 19.30
4 94647 51.9 11.5 0.754 20.55 18.80
8 94649 92.9 20.9 0.875 22.63 21.37
16 94643 159.0 35.2 0.923 27.15 25.61
32 94644 243.2 54.4 1.022 34.71 32.89
64 94648 396.6 88.0 2.264 40.94 40.11
  • 已发布评测: 2026-02-13 为 N/A(基线日期没有该配方的评测数据入库)。该配方最近的已发布 gsm8k 数据日期为 2026-04-23(运行 24816108884,TEP 并发 64 em_strict 0.9545 / em_flexible 0.9553,DEP 并发 32 em_strict 0.9575 / em_flexible 0.9583),不属于冻结的基线。

🤖 Generated with Claude Code

Bump the master image for the H100 DeepSeek-R1 FP8 Dynamo-SGLang
disaggregated MTP family from lmsysorg/sglang:v0.5.8-cu130 to
lmsysorg/sglang:v0.5.19-cu130 (Docker Hub digest
sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,
sglang tag commit 0bcd822377da7b5718e674eaf9c870d349424dd1). Model,
topology, speculative decoding, workloads and recipe references are
unchanged.

将 H100 DeepSeek-R1 FP8 Dynamo-SGLang 分离式 MTP 配方的主配置镜像从
lmsysorg/sglang:v0.5.8-cu130 升级到 lmsysorg/sglang:v0.5.19-cu130
(Docker Hub 摘要
sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,
sglang 标签提交 0bcd822377da7b5718e674eaf9c870d349424dd1)。模型、拓扑、
投机解码、工作负载与配方引用均保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Initial attempt

  • Image / SHA: lmsysorg/sglang:v0.5.19-cu130 (Docker Hub digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, tag commit sgl-project/sglang@0bcd822) at PR head ba8660e54fbebacfbbd7274e1f704317cd2e8d82
  • Change: master image only (configs/nvidia-master.yaml, key dsr1-fp8-h100-dynamo-sglang). Both CONFIG_FILE references point at upstream NVIDIA/srt-slurm recipes, so there is no in-repo srt-slurm YAML to update. Capacity checks for the single target cluster passed before the edit and before the branch was created.
  • Targeted run: none dispatched. The image bump cannot take effect on this family's launch path:
    • runners/launch_h100-dgxc-slurm.sh hardcodes the multinode dynamo-sglang container to the staged squash file /mnt/nfs/lustre/containers/lmsysorg_sglang_v0.5.8.post1-cu130.sqsh with the fixed key lmsysorg/sglang:v0.5.8-cu130, independent of $IMAGE. Only the dynamo-trt branch derives its squash file from $IMAGE, and the multinode path has no enroot import step.
    • Both upstream recipes (h100-fp8-1p1d-max-tp-mtp.yaml, h100-fp8-1p1d-max-dep-mtp.yaml) declare model.container: "lmsysorg/sglang:v0.5.8-cu130", and srtctl resolves that key through the generated srtslurm.yaml container map (src/srtctl/core/config.py), which lands on the v0.5.8.post1 squash file.
    • A sweep on this head would therefore serve the old v0.5.8.post1-cu130 container while labeling every result lmsysorg/sglang:v0.5.19-cu130. Dispatching the trimmed smoke (two 4-node deployments plus their evals) would spend H100 time to produce mislabeled data, so no e2e-tests.yml run was started.
    • Unblocking needs changes outside the Klaud edit scope: the shared launcher (also used by dsr1-fp8-h100-dynamo-trt) must derive the dynamo-sglang squash file and container key from $IMAGE, as runners/launch_h200-dgxc-slurm.sh already does; a v0.5.19-cu130 squash file must be staged on the cluster or imported on demand; and the upstream recipe's model.container must be overridden or updated. This is a launch-path limitation, not an image incompatibility finding.
  • Baseline: frozen in the PR body (published 2026-02-13, producer run 21973157671). The body also records the label mismatch: published rows say lmsysorg/sglang:v0.5.8-cu130 while the launcher's staged squash file is v0.5.8.post1-cu130.
  • Upstream source comparison: v0.5.8 @ 0189f41 (2026-01-23) → v0.5.19 @ 0bcd822 (2026-09-04). v0.5.19 and v0.5.19-cu130 share one Docker Hub digest pushed 2026-09-04; [Klaud Cold] Update dsr1-fp8-h200-sglang-mtp SGLang image to v0.5.19-cu130 / 将 dsr1-fp8-h200-sglang-mtp 的 SGLang 镜像升级至 v0.5.19-cu130 #2955 confirmed the same digest's ai.sglang.build.commit label equals the tag commit. Every server argument used by the two recipes (tp-size, dp-size, ep-size, enable-dp-attention, attention-backend flashinfer, disable-radix-cache, max-running-requests, disaggregation-mode, disaggregation-bootstrap-port, disaggregation-transfer-backend nixl, mem-fraction-static, max-prefill-tokens, chunked-prefill-size, load-balance-method, speculative-algorithm EAGLE, speculative-num-steps, speculative-eagle-topk, speculative-num-draft-tokens, stream-interval, cuda-graph-max-bs, skip-tokenizer-init, trust-remote-code, watchdog-timeout, and the launcher-injected dist-timeout) is still defined in v0.5.19 (server_args.py plus the arg_groups/ hooks), and nixl remains an accepted transfer backend. SGLANG_ENABLE_SPEC_V2=1, set in both recipes, is now a removal warning in speculative_hook.py because spec V2 is always on since v0.5.13 (#25464). SGLANG_JIT_DEEPGEMM_FAST_WARMUP was not audited. As recorded in [Klaud Cold] Update dsr1-fp8-h200-sglang-mtp SGLang image to v0.5.19-cu130 / 将 dsr1-fp8-h200-sglang-mtp 的 SGLang 镜像升级至 v0.5.19-cu130 #2955, SM90 FP8 GEMM kernels changed between these releases (CUTLASS FP8 blockwise removal in v0.5.16, #30438; SM90 FP8 decode routing fix in v0.5.19, #37018), so H100 numbers would not be kernel-for-kernel comparable once the image actually runs.
  • Coupled serving dependencies: the launcher clones NVIDIA/srt-slurm branch sa-submission-q2-2026 unpinned (current head deb1dfd9, 2026-06-17); the master router: dynamo-router 0.8.0 entry is unchanged.
  • Result: benchmark and eval deltas N/A (no updated-image run). Repairs used: 0 of 5.
  • Next step: stop without GPU use. The completion step closes this draft and retains the branch so the same candidate is not re-selected until the H100 launcher can consume the master image; a newer SGLang release will still generate a new candidate for this family.

初始尝试

  • 镜像 / SHA: lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,标签提交 sgl-project/sglang@0bcd822),PR head ba8660e54fbebacfbbd7274e1f704317cd2e8d82
  • 改动: 仅主配置镜像(configs/nvidia-master.yaml 中的 dsr1-fp8-h100-dynamo-sglang)。两个 CONFIG_FILE 均指向上游 NVIDIA/srt-slurm 配方,仓库内没有可更新的 srt-slurm YAML。编辑前与创建分支前,唯一目标集群的容量检查均已通过。
  • 定向运行: 未触发。镜像升级在该配方的启动路径上无法生效:
    • runners/launch_h100-dgxc-slurm.sh 将多节点 dynamo-sglang 容器硬编码为预置的 squash 文件 /mnt/nfs/lustre/containers/lmsysorg_sglang_v0.5.8.post1-cu130.sqsh,并使用固定键 lmsysorg/sglang:v0.5.8-cu130,与 $IMAGE 无关。只有 dynamo-trt 分支根据 $IMAGE 推导 squash 文件,多节点路径也没有 enroot import 步骤。
    • 两个上游配方(h100-fp8-1p1d-max-tp-mtp.yamlh100-fp8-1p1d-max-dep-mtp.yaml)都声明 model.container: "lmsysorg/sglang:v0.5.8-cu130",srtctl 通过生成的 srtslurm.yaml 容器映射(src/srtctl/core/config.py)解析该键,最终落在 v0.5.8.post1 的 squash 文件上。
    • 因此在该 head 上运行 sweep 会实际使用旧的 v0.5.8.post1-cu130 容器,却把所有结果标注为 lmsysorg/sglang:v0.5.19-cu130。触发裁剪后的冒烟(两个 4 节点部署及其评测)只会消耗 H100 资源并产出标注错误的数据,所以没有启动任何 e2e-tests.yml 运行。
    • 解除阻塞需要超出 Klaud 编辑范围的改动:共享启动脚本(dsr1-fp8-h100-dynamo-trt 也在使用)需像 runners/launch_h200-dgxc-slurm.sh 那样根据 $IMAGE 推导 dynamo-sglang 的 squash 文件与容器键;集群上需预置或按需导入 v0.5.19-cu130 的 squash 文件;上游配方的 model.container 需被覆盖或更新。这是启动路径的限制,不是镜像不兼容的结论。
  • 基线: 已在 PR 正文中冻结(发布日期 2026-02-13,数据来源运行 21973157671)。正文同时记录了标注不一致:已发布数据行标注 lmsysorg/sglang:v0.5.8-cu130,而启动脚本预置的 squash 文件是 v0.5.8.post1-cu130
  • 上游源码对比: v0.5.8 @ 0189f41(2026-01-23)→ v0.5.19 @ 0bcd822(2026-09-04)。v0.5.19v0.5.19-cu130 在 Docker Hub 上共享同一摘要(2026-09-04 推送);[Klaud Cold] Update dsr1-fp8-h200-sglang-mtp SGLang image to v0.5.19-cu130 / 将 dsr1-fp8-h200-sglang-mtp 的 SGLang 镜像升级至 v0.5.19-cu130 #2955 已确认同一摘要的 ai.sglang.build.commit 标签等于该标签提交。两个配方使用的全部服务参数(tp-sizedp-sizeep-sizeenable-dp-attentionattention-backend flashinferdisable-radix-cachemax-running-requestsdisaggregation-modedisaggregation-bootstrap-portdisaggregation-transfer-backend nixlmem-fraction-staticmax-prefill-tokenschunked-prefill-sizeload-balance-methodspeculative-algorithm EAGLEspeculative-num-stepsspeculative-eagle-topkspeculative-num-draft-tokensstream-intervalcuda-graph-max-bsskip-tokenizer-inittrust-remote-codewatchdog-timeout 以及启动脚本注入的 dist-timeout)在 v0.5.19 中仍有定义(server_args.pyarg_groups/ 钩子),nixl 仍是受支持的传输后端。两个配方设置的 SGLANG_ENABLE_SPEC_V2=1 现在只会在 speculative_hook.py 中触发移除提示,因为自 v0.5.13 起 spec V2 始终开启(#25464)。SGLANG_JIT_DEEPGEMM_FAST_WARMUP 未做审计。如 [Klaud Cold] Update dsr1-fp8-h200-sglang-mtp SGLang image to v0.5.19-cu130 / 将 dsr1-fp8-h200-sglang-mtp 的 SGLang 镜像升级至 v0.5.19-cu130 #2955 所记录,两个版本之间 SM90 FP8 GEMM 内核有变化(v0.5.16 删除 CUTLASS FP8 blockwise,#30438;v0.5.19 加入 SM90 FP8 解码路由修复,#37018),因此镜像真正运行后,H100 数据不会与基线逐内核可比。
  • 耦合的服务依赖: 启动脚本以未固定版本的方式克隆 NVIDIA/srt-slurm 的 sa-submission-q2-2026 分支(当前 head deb1dfd9,2026-06-17);主配置中的 router: dynamo-router 0.8.0 未改动。
  • 结果: 基准与评测变化为 N/A(没有新镜像的运行)。已用修复次数:0 / 5。
  • 下一步: 不使用 GPU 直接停止。完成步骤会关闭该草稿并保留分支,避免在 H100 启动脚本能够使用主配置镜像之前重复选中同一候选;更新的 SGLang 发布版仍会为该配方生成新的候选。

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: failed. Finishing cleanup; owned child runs will be stopped and checked before closure.


Klaud Cold:failed。正在完成清理;将先停止并确认自有子运行的状态,再关闭 PR。

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: failed. All owned runs are terminal. Repairs: 0. Runs: —.

PR closed; the exact-candidate branch is retained for manual review. The interruption does not prove image incompatibility.


Klaud Cold:failed。所有自有运行均已结束。修复次数:0。运行:—。

PR 已关闭;保留该候选的分支,等待人工审查。运行中断不能证明镜像不兼容。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

1 participant